Skip to content

Whole-spine parity for the UK: re-pinned incumbent, comparison instruments, signed register, and the E10 evidence (#686) - #747

Merged
juaristi22 merged 28 commits into
mainfrom
uk-spine-assembly-686
Aug 25, 2026
Merged

Whole-spine parity for the UK: re-pinned incumbent, comparison instruments, signed register, and the E10 evidence (#686)#747
juaristi22 merged 28 commits into
mainfrom
uk-spine-assembly-686

Conversation

@juaristi22

@juaristi22 juaristi22 commented Aug 22, 2026

Copy link
Copy Markdown
Collaborator

Whole-spine parity instruments and evidence for #686 (WS-E E10, epic #665). This PR delivers the comparison layer that decides whether the microcosm-built spine can replace the incumbent, plus two fixes the comparison itself surfaced.

Instruments

  • Incumbent re-pin 1.56.14 → 1.56.16. The pinned enhanced_frs_2024_25.h5 carried uk-data#461 (raw benunit table no longer sernum-ordered → every benunit-level variable misassigned). The parity screen compares nonzero shares, which are permutation-invariant, so it is structurally blind to this defect — is_married reads a byte-identical 0.256587 on both artifacts while every value sits on the wrong benefit unit. Re-pinned to 1.56.16 (adjudicated; identity corroborated three ways), all instruments and the four gate-battery mirrors re-cut. Reference-side movement ≤0.0046 per column, so the ±0.02 screen is undisturbed. The spine itself was never exposed: frs_spine has sorted the benunit table by id since its original ingest commit.
  • Scottish water fix (closes the WS-C deferred adjudications and the reconciliation layer #736 item 13 adjudication). FRS 2024-25 retired CWATAMT/CSEWAMT; the incumbent's fill-after-addition NaN-zeroes the charge for all 1,663 Scottish households (the whole +0.1009 divergence — incumbent-side, reported upstream as policyengine-uk-data#467). Our own level was also short (CWATAMTD is water-only); both the spine mapping and the council-tax netting now share one helper adding CSEWAMT1 at the household's own observed discount factor. Verified at every rung: share unchanged at 0.878377, Scottish level ~£395/household vs ~£185 before.
  • Signed-differences register (uk/spine_swap_signed_differences.json + loader): the migration-scoped record of adjudicated spine-vs-incumbent differences, seeded with the two water entries. Strict vocabulary, unique ids, precise per-surface scoping, no expiry (permanent adjudications; expiring suppressions stay in the per-gate exclusion registers). Advisory, not a gate — the standing health checks are the existing battery (support bounds, tail concentration, input-mass parity, aggregate-admin anchors, take-up signal, enum domains), per the US precedent; this register documents the one-time swap decision.
  • tools/verify_uk_spine_parity.py — diffs a candidate extraction against the committed reference on entity counts, nonzero shares, and (optionally, licensed) weighted totals; every difference must match a register entry or the verdict is defect. Fenced against manufactured verdicts: the reference side is never derived from the candidate, a candidate claiming the incumbent's sha is refused, and --strict fails unused register entries.
  • build_uk_efrs_parity_reference.py --candidate-h5 — extracts the same-shape surface from a candidate spine with the same producer (same engine, aliases, rounding), structurally unable to write the committed reference.
  • compare_uk_h5_payload.py --structure-only — additive verdict mode for the eventual control-vs-candidate comparison: structural predicates stay strict, value differences must be signed.
  • Identity ladder completed and repaired. New --check e7 receipts the support-channel layer (the one stacking increment that had no receipt); a negative control proves it non-vacuous. e4/e5/e6 had gone stale against E8's stacking — three distinct mechanisms (identity-keyed draws on copied rows; a population-dependent regional-uprating mean; the NHS budget normalization needing stage-time weights) — fixed with scoping driven by the artifact's own stacking flags and mass factors read from the declared operations. Green on both the post-E8 and pre-E8 spines.
  • Four in-kind columns ported (free_school_meals, free_school_fruit_veg, healthy_start_vouchers, free_school_breakfasts): a raw-mapping omission from the merged E2/E3 port, previously misattributed to E9. Coverage is now 145/145; the three contract columns land within 0.0001 of the reference.

Evidence (licensed, data/ukds/acceptance/686-spine-swap/; receipts in experiments/686-uk-spine-swap-receipts.md)

  • L0: smoke/dev/full rungs clean, every attempt landing a Logbook row; twin full builds payload-identical; record-count identity exact at 52,846 = (16,288 + 10,000) × 2 + 270.
  • L1: e4–e8 identity receipts all green.
  • L2: 145 columns compared, 0 missing, 26 beyond ±0.02 — every one attributed to an established method class (attribution follows rewrites, not just produces), with the only raw-mapping divergence (water) already signed.
  • L3: 177 input-mass totals + 47 QRF tail grids measured at spine grain (the future arming thresholds).
  • Truth columns measured through each stage's own committed cleaning function over its own pinned donor tab, on the survey-weighted basis — the convention that reproduces the E6 acceptance receipt's education figure (0.2546) exactly. E5 closer on 5/5 benchmarked shares and 4/5 levels; E7 closer on all three against their three separate truths; E6 closer on 9/13 LCFS shares and 12/13 levels, plus both ETB columns on both surfaces.

Comparison ledger (adjudication packet): committed in this branch as experiments/686-uk-spine-comparison-ledger.md — GitHub renders it directly. An interactive, continuously updated version lives as a Claude artifact (private by default; ask María for access). The committed rendition carries the UKDS EUL clause 11–12 citations.

The queue is signed

verify_uk_spine_parity --strict now returns signed_parity: 26 beyond-band divergences, all signed; 0 unsigned; no strict failure; household count exact.

The register grows from 2 entries to 13, and how they are split is the substance. No entry covers both columns where the spine is closer to its donor and columns where the incumbent is — so the LCFS class is three entries, not one: nine columns where the regime-gated draw lands closer to the donor, petrol_spending/diesel_spending where the has_fuel gate under-places incidence and the entry says so, and transport_consumption/alcohol_and_tobacco_consumption where the incumbent is closer on share and the spine on level. The four incumbent-closer columns each carry that direction in their own entry text, so re-opening one does not require re-opening the class.

Three things the re-measurement changed.

  1. The ETB rows were never undecidable. The weight basis was not the problem — the frame was. The services stage cleans one year on complete cases over a 13-column subset (4,199 rows); the incumbent cleans an 18-column subset, so a "donor share" off its frame has a different denominator. On the stage's own frame the donor share is 0.2546 unweighted — the acceptance receipt's figure to four decimals. With that fixed, the incumbent's dfe_education_spending is degenerate: 14 nonzero households in 52,846, £2 per household against a donor £3,461. Signed as a defect fix on the incumbent side.

  2. A reversal that was itself wrong. An unweighted re-measurement of the wealth columns briefly appeared to overturn the standing 2026-08-19 E5 adjudication. It did not — WAS oversamples wealth-holders by design, so its unweighted frame is not a population. On the weighted basis the original reading reproduces exactly, and corporate_wealth gains a benchmark it previously lacked. Both near-misses had one cause, recorded as a standing lesson in R5: a donor truth is a property of the frame and weighting the stage itself uses, not of the donor file.

  3. The incumbent side was measured on the wrong artifact, caught by a cross-check of every register-quoted number against its measured source. The only full incumbent H5 on disk is the 1.56.14 one staged during Retarget the UK build to FRS 2024-25: re-pin raw vintages, regenerate parity instruments, re-measure gate baselines #723, not the pinned 1.56.16 one; the pinned artifact is now fetched and digest-verified before measuring. The two differ by at most 0.0045 on any share — enough to flip alcohol_and_tobacco_consumption, which moved out of the donor-faithful entry into a two-column entry with transport_consumption that shares its evidence shape exactly. Level ratios moved in the third digit; no other verdict changed. The instrument was never wrong here — it reads the committed reference, which was always at the correct pin — the hand measurements beside it were.

One instrument change reviewers should look at

The share surface is now held to the ±0.02 #723 acceptance band rather than the reference's 6-decimal grain. Ninety further columns sit inside it on third-decimal drift from re-running every stochastic stage; signing those would have been exactly the blanket amnesty the --strict fence exists to prevent — 90 permanent adjudications describing noise, each blanket-covering its column against any future real regression.

The band governs only which magnitudes require an adjudication. Every difference down to the 6-decimal grain is still in the receipt under within_band with the in-band maximum (0.019) alongside; --share-band 0 restores the exact check; and structural differences are outside the band's reach at any band — a column appearing or vanishing, and every entity count, still signs exactly. There is a test for that specifically, run at --share-band 0.9.

--strict also now separates an entry that matched nothing on a surface this run compared from one whose surface was never examined. The weighted-totals surface stays unexamined until there is a calibrated candidate — comparing our design weights to the incumbent's calibrated ones would flag all 131 columns, the same before-medicine/after-medicine error the UC check documents — so the water-level entry reports as dormant, and a test pins that dormancy stops being available the moment a run supplies the sidecars.

Open, and deliberately not decided here

  • The incumbent's ETB per-head division — it stores a per-head figure in a household-entity column. A second upstream defect of the UK: banded CGT gains targets (by size of gain) to constrain the distribution, not just count + total #467 family, observed but not filed; it does not reconcile the levels, so it is recorded as an observation rather than the explanation.
  • communication_consumption is the one LCFS column where the incumbent is closer on level while we are closer on share. Signed inside the donor-faithful class; scoping it out is a one-line register change.
  • student_loan_balance has no like-for-like donor benchmark — signed on the standing E5 adjudication with its direction unevidenced, stated rather than papered over.

Rebased, and what follows

Rebased onto main (239 commits). The only derived value that moved was the UK spec bundle digest, because main touched uk/target_references.json and uk/target_reference_membership.json.

Spine-lane follow-up is now recorded on #757 (rescoped 2026-08-25 against the #623 calibration-seam plan): the spine defect batch the first armed calibration campaign surfaced (structural-NaN SPI columns — which the seam's NaN fence will refuse until fixed — the 85+ age-tail source stage, sidecar self-pin, support-channel gaps), per-step gating of the spine build per the US #327 doctrine with the required_at_build re-point, and the W9–W11 retirement with the markdown cleanup and the candidate-name move to microcosm_uk_2024. The calibration-only driver itself moved out to #623 (PR #743): the new seam calibrates whichever h5 it is given and never modifies a data variable, and the legacy June path retires here once the seam is proven. #757 is explicitly gated on this PR being reviewed and merged.

Refs #686, #736, #145, #665. Upstream: policyengine-uk-data#467.

🤖 Generated with Claude Code

MaxGhenis added a commit that referenced this pull request Aug 23, 2026
…_VERSION

#747 records the uk-data release tag beside the byte identity
(SOURCE_VERSION / source.version). A regeneration for another release must
not inherit the committed tag, so the identity now carries the tag (from
--release, or --version with explicit pins), the in-memory patch sets it, and
the on-disk move is anchored to the SOURCE_VERSION assignment — never a
global replacement of a tag literal, which also appears in unrelated pins
such as the registry-parity pinned_version. The lockstep test binds the
committed reference's source.version to the tool's constant once it exists.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
juaristi22 added a commit that referenced this pull request Aug 23, 2026
An armed UK national build crashed before its first stage: _source_pins
stored the full Ledger provenance block under the 'ledger_facts' role, but
role_pins_digest requires each role to carry exactly sha256 and size_bytes,
so every armed run raised ValueError at input_pins_digest.

The two halves landed independently — the strict pin contract with Logbook
adoption (#666) and the Ledger role pin with the compile-parity wiring — and
neither PR branch fires it alone, because only an armed run (--ledger-facts)
reaches this path and the licensed runs were held. It reproduces on both
#743 and #747 branches and on main.

The pin now carries the feed's verified digest and byte size; the richer
Ledger identity block already travels in safe_artifacts, source_vintages,
and the diagnostics build block, so nothing is lost.

This fix belongs upstream in #743 (the run-readiness PR whose runbook
documents the armed command); it rides the build branch until then.
juaristi22 added a commit that referenced this pull request Aug 23, 2026
Calibrates a WS-E spine at national level with the #622/#623 seam, without
the certified-input replay: the spine already carries the SPI and CGT
derivations as declarative source stages (E7/E8 ported the same pinned
sources, QRF stages, reviewed fences, and conservation receipts), so
re-running the June-convention replay would discard the spine's certified
derivation and substitute an equivalent one.

Reuses the harness pieces that are posture-independent: sha-pinned inputs,
Ledger register compilation at the declared calibration year, the frozen
national doctrine, the validated H5 writer, US-format v6 diagnostics with
the uk block, and a Logbook spool row. Scores against the incumbent on the
same frozen register before the diagnostics are written, so the score block
rides the diagnostics in a single pass (no two-pass sha dance).

Assessment posture, stated in the build record: non-certified input, the
declared battery is not enforced; the numeric fences that are honestly
measurable here (target fit, weight ratio, ESS, zero-weight strata, admin
anchors) are evaluated against the thresholds declared in uk/gates.json,
and every non-evaluated entry is listed with its reason. The spine's twelve
structural NaN columns (SPI-channel concepts undefined on the base-FRS
channel) are zero-filled from a fixed allowlist with a receipt; any NaN
outside the allowlist aborts.

Adjudication (Maria, 2026-08-23): the June convention is transitional and
retires with the completed swap; gate integration for the spine posture is
follow-up work on the #747 lane.
@juaristi22
juaristi22 force-pushed the uk-spine-assembly-686 branch from 5a546c1 to 8b893a1 Compare August 24, 2026 09:36
@juaristi22
juaristi22 marked this pull request as ready for review August 24, 2026 10:06
@juaristi22

juaristi22 commented Aug 24, 2026

Copy link
Copy Markdown
Collaborator Author

How the spine is built

The counterpart to the status diagram on #665 — same style, different question. That one maps how far the migration has got; this one maps what the spine is made of: which source each column comes from, and which kind of process puts it there. Nodes are coloured by mechanism, not by progress.

It stops at the artifact. Calibration is deliberately outside the frame — everything here produces a file carrying design grossing weights, and the solve that replaces them is downstream.

flowchart LR
  subgraph P1["1 · Raw FRS — transcription"]
    FRS(["<b>FRS 2024-25</b><br/>UKDS SN 9563 · 14 tables<br/><b>16,288 households</b>"])
    SPINE0["<b>frs_spine</b><br/>entities · IDs · weights<br/>demographics · incomes<br/>85 columns"]
    DERIV["<b>+ 6 derived stages</b><br/>employment · council tax<br/>disability · education<br/>legacy proxies · grant split"]
  end

  subgraph P2["2 · Seeded draws"]
    BRMA(["<b>BRMA counts</b><br/>public aggregate"])
    DRAWS["<b>frs_take_up · person_draws<br/>household_draws · frs_brma</b><br/>identity-keyed RNG<br/>17 columns"]
  end

  subgraph P3["3 · Wealth"]
    WASD(["<b>WAS round 8</b><br/>UKDS SN 7215<br/>15,128-row donor"])
    WEALTH["<b>was_wealth</b><br/>QRF · correlated rank draw<br/>13 columns"]
    UPRATE["<b>regional_property_uprating</b><br/>rewrites property values"]
  end

  subgraph P4["4 · Consumption"]
    LCFSD(["<b>LCFS 2023-24</b><br/>UKDS SN 9468<br/>4,202-row donor"])
    ETBD(["<b>ETB 1977-2024</b><br/>UKDS SN 8856<br/>4,199-row 2023 donor"])
    ANCH4(["<b>Anchors</b><br/>NEED energy · NHS by age/sex<br/>VAT + fare-index parameters"])
    CONS["<b>lcfs_consumption</b><br/>regime-gated QRF<br/>+ NEED raking · 19 columns"]
    ETBS["<b>etb_vat · etb_services</b><br/>QRF + NHS usage<br/>11 columns"]
  end

  subgraph P5["5 · Income — SPI channel"]
    SPID(["<b>SPI 2022-23 PUT</b><br/>UKDS SN 9422<br/>HMRC taxpayer donor"])
    ODS(["<b>HMRC published<br/>income tables</b>"])
    LEAVES["<b>frs_hmrc_spine_leaves</b><br/>FRS-side derives · 6 columns"]
    STACK["<b>spi_support_channel</b><br/><i>stacks 10,000 synthetic<br/>zero-weight households</i>"]
    SPIINC["<b>hmrc_spi_income_spine</b><br/>QRF redraw on the channel<br/>14 columns"]
  end

  subgraph P6["6 · Capital gains + derived"]
    CGTREF(["<b>HMRC CGT 2.1a + table 3</b><br/>Advani–Summers distribution<br/>SLC stocks · salsac anchor"])
    CLONE["<b>cgt_incidence_clone</b><br/><i>duplicates every household ×2</i><br/>mass split 0.5"]
    DONOR["<b>cgt_band_donors</b><br/><i>adds 270 top-tail rows</i>"]
    GAINS["<b>hmrc_cgt_gains_spine</b><br/>gain amounts by size band"]
    SALSAC["<b>salary_sacrifice · student_loans</b><br/>conversion depth · plan cohorts"]
  end

  subgraph P7["7 · Artifact"]
    ART["<b>microcosm_uk_2024 spine</b><br/>52,846 hh · 61,211 bu · 113,649 p<br/>211 declared outputs<br/><b>design grossing weights</b>"]
    IDENT["<b>Row identity closes</b><br/>16,288 + 10,000 = 26,288<br/>× 2 = 52,576<br/>+ 270 = <b>52,846</b>"]
  end

  subgraph P8["8 · What proves it"]
    L0["<b>L0</b> determinism<br/>twins payload-identical"]
    L1["<b>L1</b> identity receipts<br/>e4–e8, order-invariant"]
    L2["<b>L2</b> whole-spine parity<br/><b>signed_parity</b> · 0 unsigned"]
  end

  CAL["<b>Calibration</b> · #623<br/>replaces the weights<br/><i>out of frame</i>"]

  FRS --> SPINE0 --> DERIV --> DRAWS
  BRMA --> DRAWS
  DRAWS --> WEALTH
  WASD --> WEALTH --> UPRATE --> CONS
  LCFSD --> CONS
  WASD -.->|fuel bridge| CONS
  ANCH4 --> CONS
  CONS --> ETBS
  ETBD --> ETBS
  ANCH4 --> ETBS
  ETBS --> LEAVES --> STACK --> SPIINC
  SPID --> SPIINC
  ODS --> SPIINC
  SPIINC --> CLONE --> DONOR --> GAINS --> SALSAC
  CGTREF --> DONOR
  CGTREF --> GAINS
  CGTREF --> SALSAC
  SALSAC --> ART --> IDENT
  ART --> L0
  ART --> L1
  ART --> L2
  L2 -.->|only after review| CAL

  classDef source fill:#33415c,color:#fff
  classDef mapped fill:#2f6f52,color:#fff
  classDef drawn fill:#6b4c9a,color:#fff
  classDef imputed fill:#b8721a,color:#fff
  classDef rowop fill:#8b3a3a,color:#fff
  classDef artifact fill:#1f3a5f,color:#fff
  classDef out fill:#eee,color:#777,stroke-dasharray:5

  class FRS,WASD,LCFSD,ETBD,SPID,ODS,BRMA,ANCH4,CGTREF source
  class SPINE0,DERIV,LEAVES mapped
  class DRAWS drawn
  class WEALTH,UPRATE,CONS,ETBS,SPIINC,GAINS,SALSAC imputed
  class STACK,CLONE,DONOR rowop
  class ART,IDENT,L0,L1,L2 artifact
  class CAL out
Loading
colour mechanism what it means when a number disagrees with the incumbent
navy, rounded Source — licensed microdata or a pinned public reference ground truth. The spine is argued against these, not against the incumbent
green Transcription — read off a raw tape, no model a divergence is a defect on one side, and has to be chased to the tape. This is the signature that caught the Scottish water bug
purple Seeded draw — identity-keyed RNG reproducible for the same row, never row-for-row equal to another build. Divergence is expected and unfalsifiable by comparison alone
orange Donor imputation — QRF fitted on a donor survey two different estimators. The donor decides which side is right; the incumbent is not the referee
red Row operation — changes the row count the reason the file is 3.2× the FRS. Nothing here creates information; it reshapes who carries it
dark blue Artifact and proof what comes out, and what is checked before anything consumes it

Three processes change the row count, and they are the entire difference between 16,288 and 52,846. spi_support_channel stacks 10,000 synthetic zero-weight households so HMRC income can be modelled on a taxpayer-shaped population; cgt_incidence_clone duplicates every household with the mass split in half, so capital-gains incidence can vary within an otherwise identical household; cgt_band_donors adds 270 rows to carry a top tail no survey samples. That the arithmetic closes exactly on 52,846 is one of the strongest checks in the build — it proves the three stacking layers compose as declared rather than drifting past each other, and it is why the household count matching the incumbent exactly is meaningful while the person and benunit counts differ by +32 and −12.

Reading the colours is the point of the diagram. A green node that disagrees with the incumbent is a bug on one side. An orange node that disagrees is two estimators, settled against the donor survey. Getting that attribution wrong was the single largest source of false findings during this increment: savings_interest_income and tax_free_savings_income originate green in frs_spine and are then rewritten orange in hmrc_spi_income_spine, so attributing them to the stage that produced them reported two raw-mapping defects that do not exist. Attribution follows the last stage to produce or rewrite, which is why the diagram is a chain rather than a fan.

Why calibration is out of frame. Every weight in this artifact is a design grossing weight straight off the FRS. The incumbent's weights are calibrated, so any weighted comparison between the two compares before-medicine to after-medicine — which is exactly why the parity instrument compares unweighted shares and record counts, why the weighted-totals surface stays dormant until there is a calibrated candidate, and why the spine's UC caseload looks half the incumbent's while actually handing the solve 9% more UC-positive benunits.

@vahid-ahmadi

Copy link
Copy Markdown
Contributor

Automated review pass (Claude Code) — two passes over the diff, one low-effort and one high-effort. No verification runs; findings are from reading the diff. Ordered by what I think should block the merge.

Blocking — these change what the proofs mean

1. tools/compare_uk_h5_payload.py:~337 (apply_structure_only_verdict) — the two consumers of the signed register speak different surface vocabularies. It looks entries up with surface="payload_column" / "root_attr", but all 13 entries in the new spine_swap_signed_differences.json are scoped to nonzero_shares, weighted_totals, or entity_counts. register.matching therefore returns None for every differing column, unsigned_columns collects all of them, and --structure-only — the #686 swap-acceptance verdict — can only ever exit 1 while reporting all 13 entries as unused_ids. As far as I can tell the acceptance gate this PR exists to satisfy cannot currently pass.

2. packages/microcosm-build/src/microcosm/build/uk_runtime/signed_differences.py:~135 (UKSignedDifference.covers) — expectation is validated at load and then never consulted. covers() matches on surface + column only, so an entry signed column_missing_in_reference (e.g. num-bedrooms-net-new-column, other-investment-income-net-new-column) will also silently sign an arbitrarily large value divergence in that same column via _compare_shares's differing branch, and vice versa. The register's own scope_note — "a too-broad entry would sign a real defect" — names exactly this failure mode.

3. tools/verify_uk_identity_stability.py:~508 — the E7 receipt can pass vacuously. e7_identity_receipt's inner recompute returns {} when household_is_spi_synthetic is absent, so original is empty, the mismatch loops never execute, and the receipt reports identical_under_permutation: true / matches_stored_columns: true with exit 0 on an artifact where nothing was checked. The e4/e5 paths guard this by scoping rather than by emitting an empty receipt.

Follow-ups — real, but not merge-blocking

4. tools/verify_uk_spine_parity.py:~285 — the anti-self-comparison fence can be bypassed by an omitted field. The "reference must not be derived from the candidate" check is candidate_identity.get("sha256") == reference.source.sha256, but _candidate_identity returns {} when the extraction JSON has no source mapping (and drops None-valued keys). A candidate lacking source passes vacuously. Suggest requiring the identity to be present rather than merely non-matching.

5. tools/verify_uk_spine_parity.py:~213float("inf") vs allow_nan=False. _compare_weighted_totals sets relative = float("inf") when the reference total is 0 and the candidate's is not; that value reaches report, json.dumps(..., allow_nan=False) in main() raises, and the blanket except turns a real signed/unsigned divergence into "parity verification could not be completed" with exit 2 instead of a verdict. Any new-in-candidate column with a nonzero weighted total kills the run.

6. tools/verify_uk_spine_parity.py:~365--strict false-fails on within-band signed columns. matched_ids accumulates only from counts_report, shares_report["differing"], missing_in_candidate and extra_in_candidate; entries matching a within_band column are excluded, so a register entry signing a share column whose delta lands under --share-band is reported in unused_ids and fails --strict despite having matched a real, reported difference.

7. packages/microcosm-build/src/microcosm/build/uk_runtime/frs_spine.py:~858 (scottish_water_and_sewerage_weekly) — an unasserted data claim carries the correctness. When cwatamt1 == 0 the discount falls back to 1.0, so sewerage_gross is added at full gross on a household with no observable discount factor. Correctness rests entirely on the docstring's claim that those 22 households have CSEWAMT1 == 0 in this vintage, and nothing asserts it. A vintage refresh — or an SPI/synthetic household with a sewerage cell but no water bill — silently reintroduces the gross basis the helper exists to avoid, and this flows into council_tax through the netting in frs_council_tax.py. Worth a hard assert rather than a docstring.

On the spine itself

Both passes went looking at the donor imputation stages (QRF fits, correlated rank draws, regime gating, NEED raking), the row-stacking and mass-split arithmetic, and the signed-register semantics. Essentially everything found is in the proof and acceptance machinery rather than in the imputation chain — the construction itself held up under scrutiny, and the diagram's attribution rule (last stage to produce or rewrite) made the false-positive classes easy to avoid.

That is the good news and also the substance of the review: the spine looks sound, but several of the receipts and gates that certify it can currently pass, fail, or verify nothing for reasons unrelated to whether the data is right. Findings 1–3 are that category, which is why I would fix them before merge rather than after.

On the merge bar: for this stage, per-variable parity with the eFRS incumbent doesn't seem like the right criterion — the incumbent isn't the referee for the orange donor-imputation nodes, and its calibrated weights make weighted comparison uninformative by construction. The bar that an approval here would actually be certifying is that the spine composes as declared and that the proofs are real. The first looks true; the second needs 1–3.

juaristi22 added a commit that referenced this pull request Aug 24, 2026
An armed UK national build crashed before its first stage: _source_pins
stored the full Ledger provenance block under the 'ledger_facts' role, but
role_pins_digest requires each role to carry exactly sha256 and size_bytes,
so every armed run raised ValueError at input_pins_digest.

The two halves landed independently — the strict pin contract with Logbook
adoption (#666) and the Ledger role pin with the compile-parity wiring — and
neither PR branch fires it alone, because only an armed run (--ledger-facts)
reaches this path and the licensed runs were held. It reproduces on both
#743 and #747 branches and on main.

The pin now carries the feed's verified digest and byte size; the richer
Ledger identity block already travels in safe_artifacts, source_vintages,
and the diagnostics build block, so nothing is lost.

This fix belongs upstream in #743 (the run-readiness PR whose runbook
documents the armed command); it rides the build branch until then.
@juaristi22

juaristi22 commented Aug 24, 2026

Copy link
Copy Markdown
Collaborator Author

All seven addressed in 19799f37, plus the two spine defects the first armed calibration campaign surfaced.

Blocking

1. Surface-vocabulary mismatch — confirmed, and worse than stated. Every committed entry is scoped to nonzero_shares / weighted_totals / entity_counts, so --structure-only matched nothing, collected every differing column as unsigned, and reported all 13 entries as unused_ids. The acceptance gate could not pass.

Fixed by bridging rather than by re-scoping the register: a column_differs entry on a value-bearing surface now also covers a payload_column value mismatch on the same column, because the share instrument and the payload comparator read the same adjudicated fact through different measurements. The bridge is deliberately narrow — it does not reach entity_counts, and structural expectations never cross it. A test runs the real committed register against a differing savings column and asserts was-wealth-qrf-incidence signs it, so the two consumers can't drift apart again.

2. expectation validated then never consulted — confirmed. covers() now matches on (surface, column, expectation). num-bedrooms-net-new-column no longer signs a value divergence in num_bedrooms, and a column_differs entry no longer signs a column appearing. Tests pin both directions.

3. Vacuous E7 receipt — confirmed. It now refuses when household_is_spi_synthetic is absent rather than emitting an empty receipt, refuses when the source-key columns are missing instead of silently narrowing coverage, treats a certified column absent from the store as a mismatch rather than a skip, and records columns_compared so a green result is auditable. There were no tests on that tool at all, which is how it survived — there are now.

Follow-ups

4. _candidate_identity now requires a source block with a 64-char sha256 and refuses otherwise, so the aliasing fence can't pass vacuously. 5. A zero reference total reports relative_delta: null with reference_total_zero: true instead of float("inf"), so the run reaches a verdict instead of exit 2. 6. matched_ids now accumulates from within_band too — an entry whose divergence has shrunk under the band has matched a real, reported difference; --strict is for entries matching nothing at all. 7. The water helper now raises when a household has CSEWAMT1 > 0 and CWATAMT1 == 0, so a vintage refresh that breaks the domain claim refuses at build time instead of silently paying gross sewerage into council_tax.

The two spine defects, same commit

  • Twelve structural-NaN SPI columns. The stage left full-concept income columns NaN on the FRS channel — assessment-era honesty that made the artifact unloadable in practice (the engine's validate() refuses NaN inputs), so the campaign had to zero-fill outside the build. They now carry the adjudicated stage-time zero; the auxiliary-crosswalk guard still stops the QRF mistaking the fill for measured data.
  • The FRS 80+ age pile. No age above 80 is recorded, so the 85-89 and 90+ population targets were structurally unbindable and the 80-84 band carried the whole 80+ population — a defect the incumbent shares. New age_tail declarative stage disperses the pile from a sex-specific inverse CDF over committed ONS band populations, keyed on person_source_id so clone twins agree, running last in the plan so nothing that conditions on age sees a different input.

Whole-spine parity re-verified after the register semantics changed: signed_parity, 0 unsigned, --strict clean, household count exact.

On the merge bar

Agreed, and it is the bar the PR now argues for explicitly. Per-variable parity with the incumbent is the wrong criterion here — the incumbent isn't the referee for the orange donor-imputation nodes, and its calibrated weights make weighted comparison uninformative by construction. What an approval certifies is that the spine composes as declared and that the proofs are real. Per-step gating of each imputation stage, so a regression is caught where it happens rather than at a whole-spine diff, is filed as #757.

@vahid-ahmadi

Copy link
Copy Markdown
Contributor

Follow-up verification pass over 19799f37 (Claude Code, high effort, diff only — no build or test execution). I reviewed the fixes themselves rather than re-reviewing the PR, since they land in the components that decide whether the gate passes.

Eight of the nine hold. Verified sound, not just taken on trust: the (surface, column, expectation) gating; the E7 refusals (both raise an uncaught ValueError, nothing swallowed) and columns_compared accuracy, including an absent stored column now counting as a mismatch; _candidate_identity's sha256 requirement together with the now-unguarded candidate_identity["sha256"] deref (safe — the raise precedes it); relative_delta: null having no other consumer of that key in the repo; the water helper's ~billed & (sewerage_gross > 0) guard; and the age_tail inverse CDF (searchsorted(..., side="right") with a min() clamp, low + int(draw * width) giving 80-84 / 85-89 / 90-97, ages rewritten in place so no rows or weights move and the 80+ total is conserved, both draws keyed on person_source_id so clone twins agree).

The bridge is the exception, and it needs one more turn.

1. uk_runtime/signed_differences.py:132covers() ignores scope.entities, and the bridge makes that gap load-bearing

Every entry in spine_swap_signed_differences.json is entity-scoped — scottish-water-incumbent-nan-zeroing carries entities: ["household"] on column water_and_sewerage_charges — but tools/compare_uk_h5_payload.py:354 loops for key, table in report["tables"].items() and looks up by column only. So a value mismatch on a same-named column in the person or benunit table is signed by a household-scoped adjudication.

That is a previously-unsigned divergence becoming silently signed, which is the failure mode the register exists to prevent — reintroduced by the fix for it. It also means the current signed_parity / 0 unsigned / --strict clean result can be achieved with a cross-entity divergence signed by the wrong adjudication, so I don't think that result is yet proof of what it says. The fix looks small: compare scope.entities against the table's entity inside covers(), and have the payload comparator pass the entity it is iterating.

The bridge is otherwise correctly narrow, as claimed — expectation is compared first so structural expectations cannot cross, and entity_counts is properly excluded.

2. uk_runtime/signed_differences.py:137 — an empty columns tuple would blanket-sign every payload column

The bridge reuses the surface-wide semantics (not self.columns or column in self.columns), so a column_differs entry on nonzero_shares / weighted_totals with an empty columns tuple would sign every payload column in every table. No committed entry is column-less today, but the docstring documents empty-columns as a supported form for entity_counts, so nothing prevents one being added later. Worth refusing the combination explicitly rather than relying on no one writing it.

3. tools/verify_uk_spine_parity.py:314weighted_totals still missing from matched_ids

matched_ids now accumulates from counts_report and from shares_report's differing / within_band / missing_*, but never from _compare_weighted_totals's differing. An entry adjudicated only on the weighted_totals surface therefore reads as register rot under --strict. Latent rather than live, since that surface is dormant — but it will bite exactly when the surface is un-dormanted for a calibrated candidate, which is the moment it starts mattering.


On the merge: 1 is the only thing I'd hold for. It's a narrow change and once the entity scope is honoured — with a test pinning a person-table divergence not being signed by the household-scoped entry, mirroring the register test you already added — I'm happy to approve.

juaristi22 and others added 12 commits August 25, 2026 11:01
The frozen instruments were pinned at the incumbent data package's
1.56.14, which carries uk-data#461: from the 2024-25 FRS release the
raw benunit table is no longer ordered by sernum, so every
benunit-level variable — benunit_id included — landed on the wrong
benefit unit relative to the model's sorted-id entity order.

The parity screen compares unweighted nonzero shares, which are
invariant under a row permutation, so it cannot see this. is_married is
one of the 145 columns it compares and reads a byte-identical 0.256587
on both artifacts while every value sits on a different row. Signing
whole-spine parity against 1.56.14 would have frozen the upstream
defect into the contract as though it were correct.

Verified on the new artifact before trusting it: entity counts and
column surface unchanged (113,617 / 61,223 / 52,846; 145 layers),
clone_index uniformly 0 so it is still the pre-clone artifact that
compares row-for-row with the spine grain, and benunit_id now sorted
ascending where 1.56.14 was not. The identity is corroborated three
ways — the HF LFS metadata at the tagged revision, the repo's own
releases/1.56.16/release_manifest.json, and a local hash.

Reference-side movement is small: 39 of 145 shares move, none by more
than 0.0046, so the +/-0.02 screen is undisturbed. The licensed
weighted register moves on all 128 comparable columns, but that is the
already-signed register-realization class — 1.56.15 changed the UC
caseload targets and the incumbent's calibration re-solves with
unseeded dropout — and stays far inside the gross-mass fence. No gate
threshold moves here; #686 arms the re-measured baselines later.

uk/frs_release.json is deliberately untouched: its revision names the
raw UKDS zip, a different artifact that did not change.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
FRS 2024-25 retired CWATAMT and CSEWAMT ("Final value after discount").
The headers survive but carry no data in any of the 16,288 households,
and the FRS replaced them with the Scotland-only CWATAMT1/CSEWAMT1
("Weeklyised gross annual dom. water/sew. charge on bill", "DV created
in 2024-25 as variable was removed from the dataset").

Two things were live. The incumbent adds the retired cell before
filling, so a wholly blank CSEWAMT propagates NaN and zeroes the charge
for all 1,663 Scottish households that have one; our per-column fill
left it standing, which is the whole +0.1009 share divergence flagged
on #736 — we are correct, the incumbent is defective, and the gap
reproduces on the raw tab at +0.1021 before composition. Separately our
own level was short: CWATAMTD is the water charge alone, about GBP 185
per Scottish household against roughly GBP 490 for England and Wales.

Both consumers now call one helper, so the amount netted from the
council tax bill is exactly the amount charged as water and sewerage —
an inconsistency the incumbent still has, since its netting fills per
column while its charge fills after the addition. The helper adds
CSEWAMT1 discounted at the household's own observed CWATAMTD/CWATAMT1,
which keeps the retired cells' after-discount meaning instead of
silently switching to a gross basis. The factor is well behaved: range
(1/3, 1], never above 1, for the 1,641 households with a positive gross
bill. The other two domains cannot be moved by it — the 22 with a
recorded CWATAMTD but no gross cell have zero sewerage, and the 21 with
no council-tax cells stay at zero.

The level fix does not move the nonzero share (0.8783767190569745
either way, matching the existing spine evidence to every digit), so
the share difference stands alone as an incumbent-defect divergence.

The fixtures supplied a non-missing CSEWAMT and so never exercised what
the tab contains; they now carry the retired cells blank, and a
regression test pins that a blank retired cell and an absent one give
the same non-zero answer.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Disclosure-controlled evidence for the two changes above: the pin
identity and its three-way corroboration, the structural verification
of the new artifact, the defect footprint across benunit/person/
household, the reference-side share movement, the licensed weighted
register's digests and relative drift, and the FRS 2024-25 water cell
retirement with the factor domain and level comparison.

Digests, column counts, unweighted shares and relative deltas only per
CD171 5.2.1; the licensed register itself stays uncommitted under the
UKDS EUL.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The whole-spine comparison has one rule: anything differing between the
spine and the frozen incumbent that is not signed is a defect. That rule
needs somewhere to read the signatures from, so this adds the register,
its loader, and the validation that keeps it trustworthy.

Each entry names its class, the exact surface and columns the difference
is expected to appear on, disclosure-safe magnitude evidence, the
adjudicator and the date. The loader enforces the three vocabularies,
unique kebab-case ids and ISO dates, and a test refuses a column-surface
entry that names no columns — an unscoped entry would quietly absorb
unrelated divergences, which is precisely the failure this register
exists to prevent.

It sits above the per-gate reviewed-exclusion registers rather than
replacing them. Those are per-gate, per-reference and expiring, because a
suppression must not outlive its reason; a signed difference is a
permanent adjudicated fact. So expires_on is rejected outright, with the
error pointing at the exclusion registers instead.

Seeded with the two Scottish water adjudications, deliberately scoped
apart: the incumbent's NaN-zeroing signs the share surface, the
successor-cell level change signs weighted totals and leaves the share —
which it does not move — unsigned. The E4-E8 method classes and the #723
beyond-band columns are transcribed as each is re-measured against the
re-pinned reference and scoped to what is actually observed, rather than
carried across on prose classification.

The country package declares everything it ships, so the resource is
registered there; that moves the UK spec bundle sha, re-pinned in the
same change. Wheel-checked: the resource ships beside its siblings.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The swap decision rests on this comparison, so both halves of it live
here: a candidate mode on the existing extractor, and the diff that
holds its output to the signed register.

The candidate mode reuses build_reference and build_weighted_totals
unchanged and only swaps the identity block. That is the point — if the
two sides were measured by different producers, the diff would be
between two measurement methods rather than two artifacts. A candidate
is identified by its own sha256 instead of being checked against the
incumbent pin, and the mode cannot write the committed reference: it
refuses any destination inside the country package, and --check is
refused outright because a candidate can never satisfy a check against
the incumbent pin.

verify_uk_spine_parity.py compares the record-count identity exactly,
per-column nonzero shares at the reference's own six-decimal grain with
the column-set difference both ways, and optionally the two licensed
weighted registers as relative deltas only. Every difference must match
a register entry; anything else is a defect and exits 1.

Two fences stop the verdict being manufactured. The reference side is
always the committed instrument, and a candidate extraction claiming the
incumbent's own sha256 is refused rather than compared — a copied
reference would pass by construction. --strict, the swap-acceptance
posture, also fails when a register entry matched nothing, so the
register cannot decay into a blanket amnesty as the spine changes.

Verified end to end against the E8 spine artifact: 142 columns compared,
the household record-count identity holding while the known donor-
composition deltas show on persons and benefit units, and
water_and_sewerage_charges reproducing at +0.100897 and binding to its
signed entry rather than counting as unsigned.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The swap comparison asks a different question from payload identity: the
control and candidate builds should share a surface exactly and differ in
values only where a difference has been adjudicated.

This re-verdicts the same measurements rather than relaxing them. Every
structural predicate stays strict — same keys and stored kinds, row
counts, column lists in order, dtypes, indexes, root-attribute names —
and each differing column or root attribute must name an entry in the
committed signed-differences register. A signature excuses a differing
value, never a differing surface, which is pinned by a test that supplies
a signature for an added column and still expects failure.

payload_identical is still computed and reported in both modes, so a
structure-only receipt stays comparable with a full-mode one, and
--signed-differences is refused outside the mode rather than silently
ignored.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Not an acceptance run — the E8 spine predates both the re-pin and the
water fix, and the register is deliberately seeded with only the two
adjudications made so far. It records that the instrument behaves as
specified on real artifacts before the licensed ladder depends on it:
the household record-count identity holds, the known donor-composition
deltas surface on persons and benefit units, the E9 derived-benefit
columns are still the three missing on the candidate side, the water
column binds to its signature instead of counting as unsigned, and the
weighted-totals entry correctly reports as unused when no weighted
registers were supplied.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Both UK resource tuples in test_country_spec.py enumerate the package's
resources exactly, so #686's signed-differences register has to appear in
them. Caught by the full build-shard run rather than the targeted ones.

The first test's name records the #717 question it was written to answer,
but what it does now is pin the whole legacy-JSON list; a note says so, so
the next reader is not puzzled by a resource arriving in a test that says
nothing was added.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
L0 ladder built from raw licensed tabs on this branch: smoke, dev and full
all clean, every attempt landing a Logbook row. The record-count identity
closes exactly at full scale — (16,288 + 10,000) x 2 + 270 = 52,846 — and
the Scottish water fix reproduces at every rung.

Parity against the re-pinned reference finds 26 columns beyond the band,
against 27 in the #723 screen, which is what a re-pin that moved no
reference share by more than 0.0046 should produce.

Includes a correction to how divergences are attributed. The surviving
value belongs to the last stage to produce *or rewrite* a column, and
attributing by producing stage alone manufactures false findings:
savings_interest_income and tax_free_savings_income originate in
frs_spine and are rewritten by the SPI channel, so a naive pass reports
them as raw-mapping divergences outside every signed class — the exact
signature that made the water defect real. With rewrites folded in, every
beyond-band divergence lands in an established class and the only
raw-mapping one is already signed.

The queue itself is left unadjudicated. The verdict is defect by
construction because the register holds only the two water entries, which
is the correct starting state; each class needs a ruling before it becomes
an entry, and the E9 derived-benefit gap is the one item that may argue
against swapping rather than for signing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Twin determinism passes: two independent full builds are payload-identical
across every store key, differing only in bytes because HDFStore stamps
write times.

All four identity receipts pass the property they exist to prove — that
recomputation is invariant to row order. e4, e5 and e6 nevertheless report
matches_stored_columns false, and the cause is the instrument, not the
spine. _frs_only_frame scopes the survey channel by excluding
household_is_spi_synthetic alone, which was right when the SPI channel was
the only layer stacking rows. E8 added two more: the capital-gains
incidence clone and the 270 band donors. Those rows carry values copied
from their sources, so recomputing an identity-keyed draw for a clone's own
household id disagrees with a stored value that was never drawn for that
id.

The measurement is unambiguous: excluding all three flags leaves exactly
16,288 households, the raw FRS count, which is the scope the #723 receipts
ran at and passed. e8's own receipt passes because it recomputes the clone
and donor logic explicitly instead of assuming unstacked rows, and e4's
mismatch list is entirely identity-keyed draw columns — precisely those a
clone inherits rather than draws.

Recorded as unsigned and unfixed: the scope must exclude every stacked
layer, the fix belongs with the e7 receipt work since both are ladder
maintenance, and until then these three results say nothing about the
spine and the L1 leg of the gate is not satisfied. Also noted: e5 and e6
report the failure with an empty mismatch map, which is not actionable
evidence and should name what disagreed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Building a spine that carries every stage through E8 and running the whole
ladder surfaced that e4, e5 and e6 had gone stale. Each was written when
the SPI support channel was the only stage stacking rows; E8 added the
capital-gains clone and the CGT band donors after them.

The failures are in the instrument, not the spine, and the experiment that
proves it is the unchanged tool against two artifacts: the pre-E8 spine
passes e6, the post-E8 spine fails it.

Three mechanisms, which is why one fix did not cover them. e4 recomputes
identity-keyed draws and a stacked row carries a value copied from its
source, never drawn for its own id. e5's regional uprating scales to a
per-region mean over the frame's owner households, so stacked rows move
the denominator — confirmed by running it both ways on one artifact.
e6 fails scoped as well as unscoped: its NHS allocation normalizes against
an absolute budget, so it needs stage-time weights, and it divided out the
SPI channel's share while the clone's mass_split went unrestored.

Scoping now excludes every stacked layer from one declared flag list, and
the weight divisor reads its factors from the declared operations instead
of hardcoding them.

The divisor is driven by the flags the artifact actually carries rather
than by the committed roster. The first implementation used the roster and
divided the clone factor out of a spine built before that stage existed,
skewing the comparison the other way; the pre-E8 artifact caught it at
once. Both vintages now pass.

Recorded at the helper: a new mass-redistributing op kind has to be
registered there, and the failure mode if it is not is a receipt silently
comparing against the wrong grossing scale.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Two gaps the whole-spine parity screen surfaced, both in already-merged
increments rather than in new work.

The spine never mapped healthy_start_vouchers, free_school_breakfasts,
free_school_fruit_veg or free_school_meals, so parity reported three of
them as columns the candidate does not produce. The incumbent maps them
straight off the person tapes in create_frs, which makes this an E2/E3
omission rather than deferred work — #685 is UC deduction attributes, bus
fares and WAS debt, none of which touch these. heartval is on both the
adult and child tapes and the three school columns are child-only, so
adults read zero rather than propagating NaN, which the test pins.
Measured against the reference the three contract columns land at
-0.000089, -0.000001 and +0.000008; free_school_breakfasts is not an
engine-known variable so it never enters the 145-column surface.

The identity ladder had no e7 check at all — the increment that
introduced row stacking, and so the one whose interaction with E8 broke
e4, e5 and e6, was the only one without a receipt. It now receipts the
support-channel layer: each entity's channel and clone index, the
composite source key, and the propagation of a household's channel to
its persons and benefit units. Bitwise on both surfaces, since these are
labels and integer indices. A negative control confirms it is not
vacuous: flipping one channel label fails it and names the column.

It deliberately excludes the employer_pension_contributions = 3 x
employee_pension_contributions derive. That is a real E7 layer, but E8's
salary_sacrifice rewrites the multiplicand in place afterwards and the
relation survives on only 95.9% of survey-channel persons, so asserting
it would fail for the wrong reason — the #721 rewrites-provenance class.

The coverage manifest is regenerated for the new source-manifest hash
(145 required, 0 exclusions, unchanged) and the spec bundle sha re-pinned.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
juaristi22 and others added 15 commits August 25, 2026 11:01
The living adjudication packet is a Claude artifact, which is private by
default — a reviewer clicking through from the PR hits an access wall.
The packet therefore also lives in the repo, where GitHub renders it in
the review itself and it travels with the branch.

Same content as the interactive page: what the shares measure and their
limits, the three incumbents, per-class share tables with donor or
raw-source truth, the levels table with the taxable-to-taxable SPI
comparison, coverage, diagnostics, and the open items. Adds the UKDS EUL
clause 11-12 citations and the disclosure-control statement, which apply
wherever these aggregates are posted.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
UC had no dedicated row despite being what the armed calibration binds.
Measured modeled universal_credit through the engine on both artifacts:
unweighted the spine carries 9% more UC-positive benunits than the
incumbent at an equal per-recipient level, and the reported column is
exact against the raw tab to every digit. The halved weighted caseload
is the calibration boundary itself, not a spine defect — the incumbent's
weights already embody a UC caseload target since 1.56.15, so the
comparison is before-medicine to after-medicine.

Recorded watch-item: would_claim_uc frozen at 0.55 (U8, uk-data#452),
the pre-registered lever if the armed run's UC fit is strained.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Main moved uk/target_references.json and uk/target_reference_membership.json,
both of which the UK spec bundle hashes, so the pinned bundle digest in the
country-bundle test moves with them. Re-cut against the rebased base; the
contract mirrors, the parity reference and the coverage manifest all verify
unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Takes the whole-spine parity verdict from defect to signed_parity with
nothing unsigned and no strict failure. Two changes make that honest rather
than merely green.

The register grows to thirteen entries, split so that no entry covers both
columns where the spine is closer to its donor and columns where the
incumbent is. Donor evidence is re-measured through each stage's own
committed cleaning function on the survey-weighted basis, which reproduces
the E6 acceptance receipt's education figure exactly; that settles the ETB
weight-basis question and exposes the incumbent's dfe_education_spending as
degenerate at fourteen nonzero households in 52,846.

The instrument's share surface moves from the reference's six-decimal grain
to the #723 acceptance band. Ninety columns sit inside it on third-decimal
drift, and signing those would have been the blanket amnesty the register is
built to prevent. Nothing is hidden: in-band differences are reported under
their own key with the in-band maximum, --share-band 0 restores the exact
check, and structural differences stay outside the band's reach at any band.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Re-measures every donor comparison through each stage's own committed
cleaning function on the survey-weighted basis, adds the level ratios
alongside the shares, and replaces the E6 "needs ruling" and E5
"carry-forward?" sections with what the evidence now shows.

The ETB rows are no longer undecidable: the stage cleans a thirteen-column
subset of one year and weights by hhold_adj_weight, and on that frame the
incumbent's education column is degenerate at fourteen nonzero households in
52,846. The incumbent's per-head division is recorded as a second upstream
defect, observed but not filed.

Also records the correction that a mid-review unweighted re-measurement of
the wealth columns appeared to overturn the standing E5 adjudication and did
not: WAS oversamples wealth-holders, so the weighted basis is the population
one, and on it the original reading stands.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
R5 receipts the signed_parity verdict and, more usefully, the two ways the
donor measurements had been wrong before it. Both had the same cause: a donor
truth was computed on a frame the stage does not use. Once against the ETB
services frame, which produced a false blocker and had been recorded as an
unpinnable weight basis; once against the WAS frame unweighted, which briefly
appeared to overturn a standing adjudication.

Calling each stage's own cleaning function reproduces the E6 acceptance
receipt's education figure to four decimals, which is the check that the
convention is the right one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The signing tables report population mean per household; the Levels section
reports conditional mean per carrier. They can disagree on which side is
closer, so the section now says which one the signing used and why: the
population mean is what a calibration target binds and what survives a
difference in incidence.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
An earlier pass read the only full incumbent H5 on disk, which is the 1.56.14
published artifact rather than the pinned 1.56.16 one. The two differ by at
most 0.0045 on any share, but that was enough to flip one verdict:
alcohol_and_tobacco_consumption reads as ours at 1.56.14 and as the incumbent
at the pin. It moves out of the donor-faithful entry into a two-column entry
with transport_consumption, which shares its evidence shape exactly —
incumbent closer on share, ours closer on level.

The pinned artifact is now fetched and verified by digest before measuring,
and the correction is recorded in the ledger rather than quietly applied.

Also rewrites "What the numbers are", which still claimed the ledger measures
incidence and never level. That is true of the committed instrument and is
precisely why the ledger exists; the document itself evaluates both, and the
section now says which quantity is operative when they disagree.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A cross-check of every register-quoted number against its measured source
found the incumbent side had been measured on the 1.56.14 artifact staged
during #723 rather than the pinned 1.56.16 one. The instrument itself was
never wrong, since it reads the committed reference; the hand measurements
beside it were.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The earlier pin correction updated the level ratios and the
other-residential share but left savings, property_wealth and
corporate_wealth quoting the 1.56.14 incumbent. Caught by re-running the
cross-check that compares every register-quoted number against its source.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Eleven of the thirteen evidence fragments did not match any heading in the
file they cite — two of them since before this branch, the rest because the
ledger's section titles changed when the queue was signed. A dead fragment
fails silently in a browser, so the rot only surfaces when a reviewer clicks
and lands nowhere, which is exactly when the pointer needed to work.

The new test derives GitHub's own anchor form from each cited file's headings
and requires the fragment to be among them, so renaming a section now breaks
CI instead of breaking an audit trail.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two lanes of fixes that belong together: the review found the spine's
construction held up and its proof machinery did not, and the first armed
calibration campaign found two spine defects it had to work around outside
the build. Both are cheaper to fix now than after a swap.

Spine defects. The SPI stage left twelve full-concept income columns NaN on
the FRS channel, which made the artifact unloadable in practice — the engine
refuses NaN inputs and any finiteness fence on the calibration path refuses
the frame — so the campaign zero-filled them outside the build. Zero is the
adjudicated stage-time semantics for a concept the instrument never asked
about. Separately, the licensed FRS records no age above 80, so the 85+
population targets were structurally unbindable and the 80-84 band carried
the whole 80+ population; the new age_tail stage disperses the pile from a
sex-specific inverse CDF over committed ONS band populations, keyed on
person_source_id so clone twins agree, running last so nothing that
conditions on age sees a different input.

Proof machinery. Three findings could change what a proof means: the
register and the payload comparator spoke different surface vocabularies, so
--structure-only could only ever fail; `expectation` was validated and then
never consulted, so a column's signed *appearance* also signed its values;
and the E7 receipt reported green over an empty recomputation. Four quieter
ones: a candidate omitting its source identity passed the aliasing fence
vacuously, a zero reference total produced float("inf") that the encoder
refused, --strict false-failed on within-band signed columns, and the
Scottish water discount fallback rested on an unasserted vintage claim.

Whole-spine parity re-verified after the register semantics changed:
signed_parity, 0 unsigned, --strict clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
CI caught eight failures I should have: adding age_tail moves three derived
surfaces, and I ran targeted suites instead of the shard.

The stage roster is pinned in three places, and the E8 contiguity invariant
asserted that the E8 block sat immediately before the certified pair — true
only because E8 happened to be the last increment. What the invariant is
actually protecting is that E8 stays contiguous and the certified pair stays
at [-2:], where the frozen-copy lockstep test reads it; both survive a spine
tail. The test now says that, with age_tail declared as a POST_E8 block so
the next stage after it has to make the same decision deliberately.

The release input-coverage manifest carries the source manifest's digest, so
it regenerates: 145 required columns and 0 exclusions unchanged, confirming
age_tail adds no column — it rewrites `age`, which was already required. The
stale digest was also what failed the preflight battery's
manifest-current gate.

Verified against the packaging gate this time: wheels built, installed into a
clean venv, suite run from there.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…lted

The re-review found the payload bridge made an old gap load-bearing:
covers() matched on surface and column only, so a household-scoped
adjudication could sign a same-named column on the person or benunit table
once a per-table comparator started consulting it. That is an unsigned
divergence becoming silently signed — the failure the register exists to
prevent, reintroduced by the fix for it.

Lookups now carry the entity, the payload comparator passes the store key it
is iterating, and the bridge forwards the entity scope rather than widening
it. The check earned its keep immediately: student_loan_balance was scoped to
household and is a person column.

The loader now also refuses a column-surface entry that names no columns or
no entities, so the surface-wide form survives only on entity_counts, where
the "column" is itself an entity name.

Weighted-totals matches were already counted in matched_ids before unused is
computed; a test now pins that a totals-scoped entry reads as matched rather
than as rot, since the surface is dormant and the accounting would otherwise
first be exercised on a calibrated candidate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@juaristi22
juaristi22 force-pushed the uk-spine-assembly-686 branch from 86f5574 to 76e39f9 Compare August 25, 2026 09:10
@juaristi22

Copy link
Copy Markdown
Collaborator Author

Fixed in 76e39f9b, rebased onto current main.

1. covers() ignores scope.entities — confirmed, and it caught a live mis-declaration

You're right that the bridge made this load-bearing, and right that the previous signed_parity result wasn't yet proof of what it said. covers() and matching() now take an entity, the payload comparator passes the store key it is iterating, and the bridge forwards the entity scope rather than widening it.

Honouring the scope immediately found one wrong: was-student-loan-balance-fold scoped student_loan_balance to household, and it is a person column. My entry, written while noting the column sits at a different grain from the wealth columns beside it — and then scoped to the wrong one. Harmless while entities were never consulted; with the check it stops matching, which is the correct behaviour and exactly your point. Re-scoped to person, and there's now a test asserting every committed entry's declared entities agree with the parity reference's input_entities, so a future mis-scope fails rather than sits latent.

Tests are the pair you asked for, end to end through the comparator: a person-table divergence in water_and_sewerage_charges is not signed by the household-scoped entry (exits 1, unsigned_columns == ["person.water_and_sewerage_charges"], matched_ids == []), and the same entry does sign the household-table divergence.

A caller that genuinely cannot determine an entity passes None and gets no entity filtering — documented at the call site. That is safe on the surfaces where it happens, because a Frame's column names are globally unique; the payload comparator is the one surface that iterates per table, and it always knows its entity.

2. Empty columns would blanket-sign every payload column — fixed at the loader

Rather than rely on nobody writing one, the loader now refuses an entry on a column surface (nonzero_shares, weighted_totals, payload_column) that names no columns or no entities. The surface-wide form survives only on entity_counts, where the "column" is itself an entity name and a blanket entry can't absorb anything unrelated. covers() keeps a defence-in-depth check for a register built by another path. Three tests: both refusals, and entity_counts still permitted.

3. weighted_totals missing from matched_ids — already covered, and now pinned

This one I'd push back on: matched_ids.update(...) from totals_report["differing"] is at verify_uk_spine_parity.py:370, inside the optional-totals block, and unused is computed at :381 — after it. So a totals-scoped entry that matched does count as matched today.

Your instinct that it would first bite when the surface is un-dormanted is the right worry though, and nothing was pinning it, so there's now a test: with both sidecars supplied and a totals-only entry that matches, --strict passes with matched_ids == ["totals-signed"], unused_ids == [], dormant_ids == [].


Whole-spine parity re-run with entity scoping enforced: signed_parity, 0 unsigned, --strict clean, household count exact — and it now carries the meaning you were asking for, since a cross-entity divergence can no longer be signed by the wrong adjudication.

Rebased onto origin/main (through 5abda614); the UK spec bundle digest is unchanged by main's six new commits, and the pin suites, packaging gate and parity instrument all pass on the rebased tree.

@vahid-ahmadi vahid-ahmadi left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving. 76e39f9b is green on all four checks and is the commit I reviewed.

I verified the entity-scoping fix in the diff rather than reading the write-up: covers() and matching() take an entity, and every call site now passes one — the payload comparator passes its store key, _compare_shares and _missing pass entities.get(column), and _compare_weighted_totals gained an entities parameter fed from reference.input_entities. The bridge forwards the scope rather than widening it. Refusing an unscoped column-surface entry at the loader is a stronger fix than the one I suggested, and entity_counts is correctly exempt since its "column" is the entity. The re-scoped was-student-loan-balance-fold entry has its magnitude_evidence prose updated to match rather than left stale, and the test asserting every committed entry's entities agree with the reference's input_entities is what stops the next one sitting latent.

Your push-back on finding 3 was correct and the finding was wrong. I read verify_uk_spine_parity.py at 76e39f9b: matched_ids.update(...) from totals_report["differing"] is inside the optional-totals block and unused is computed after it, so a totals-scoped entry that matched has always counted as matched. The pinning test is a good outcome from a bad finding.

That the entity check immediately caught student_loan_balance scoped to household when it is a person column is the substantive result here — it means the previous 0 unsigned was reachable with a real mis-scope in the register, and the current one is not.

Worth stating what this approval covers, since the PR is careful about that distinction elsewhere: I reviewed the diff and confirmed CI is green. I did not run the build, the parity instrument, or the suites locally, so the signed_parity / 0 unsigned / --strict clean result is one I have read the code for rather than reproduced. On that basis the spine composes as declared and the proofs now mean what they say.

@juaristi22
juaristi22 merged commit 2807e99 into main Aug 25, 2026
4 checks passed
MaxGhenis added a commit that referenced this pull request Aug 25, 2026
…the reformatted manifest and re-cut the digests over the union

Main's #747 reformatted uk/source_stages.json and moved the attested
surfaces again, so the merge takes main's file and re-applies the five
WAS-stage edits (derived split, chain order, outputs, nonnegative outputs,
notes incl. the per-segment seed sentence) in its format, regenerates
release_input_coverage_manifest.json (145 required / 0 exclusions
unchanged), re-pins the UK spec_sha256 and re-cuts the three gate-battery
digests over the union - the d70ea39 pattern, second application.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants